Computers in Biology and Medicine
○ Elsevier BV
Preprints posted in the last 7 days, ranked by how well they match Computers in Biology and Medicine's content profile, based on 128 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.
De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.
Show abstract
Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.
Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.
Show abstract
SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.
Patil, A.; Barathe, R.; Tate, D. M.; Kate, K.; Pande, S.; Gawande, N.; More, A.; Mahadik, S.; Berde, K.; Singhvi, R.
Show abstract
Introduction: Polyendocrine metabolic ovarian syndrome (PMOS), formerly known as polycystic ovary syndrome (PCOS), is a common endocrine disorder affecting women of reproductive age. Besides reproductive and metabolic disturbances, PMOS negatively impacts psychological well-being and quality of life. Despite available treatment options, there remains a need for safe and effective therapies that improve both clinical symptoms and fertility outcomes. Aim: To compare the efficacy of VAMHA and MYRHA tablet combination therapy with standard non-hormonal therapy in restoring regular menstruation. Secondary objectives included assessment of ovulation, menstrual symptoms, polycystic ovarian morphology, hormonal and metabolic parameters, anthropometric measures, and skin manifestations. Study Design: Open-label, randomized, multicentre, prospective comparative clinical study. Methods: Seventy-one women with PMOS were randomized to Group A (n=37) or Group B (n=34). Group A received VAMHA and MYRHA tablets (2 tablets each), while Group B received Metformin 500 mg plus Myoinositol 600 mg (1 tablet), twice daily for 180 days. Data were recorded in Case Report Forms. Statistical Analysis: Continuous variables were summarized using mean and standard deviation, while categorical variables were expressed as frequencies and percentages. Appropriate statistical tests, including Chi-square, were used. A p-value [≤]0.05 was considered significant. Results: Significantly more participants in Group A achieved regular menstrual cycles than Group B (31 vs. 22; p<0.05). Ovulation occurred in 16 participants in Group A compared with 6 in Group B (p<0.05). Both groups showed significant improvement in menstrual irregularity and related symptoms. Significant reductions in Anti-Mullerian Hormone (AMH), fasting insulin, and body mass index (BMI) were observed in both groups (p<0.05). Resolution of polycystic ovarian morphology occurred in 13 participants (38.23%) in Group A and 10 (33.33%) in Group B. Both treatments were well tolerated with no major safety concerns. Conclusions: VAMHA and MYRHA combination therapy was superior to standard non-hormonal therapy in improving menstrual regularity and ovulation. It also produced favourable metabolic, hormonal, and ultrasonographic outcomes, suggesting its potential as a safe and effective option for comprehensive PMOS management and fertility enhancement.
Ekambarapu, L.; Pendyal, A.; Lin, A.; Alwakeel, M.; Rajaratnam, A.
Show abstract
Background: Unstructured biomedical data, such as echocardiography reports, are rich in information but time consuming to analyze at scale. Rule-based, regular expression-driven terminology mapping can only extract individual variables while large language models (LLMs) offer scalable and clinically meaningful interpretations of heterogeneous disease processes. Right ventricular dysfunction (RVD) is an example of a multifactorial disease state in which key structural and physiologic features are captured both narratively and in structured fields, making it an ideal test case for evaluating whether LLMs can recover complex phenotypes that rules based methods routinely miss. Purpose: To compare an LLM-based extraction method to a conventional rules-based schema for identifying and phenotyping echocardiographic features associated with RVD in a large TTE dataset. Methods: MIMIC-III NOTE2NUM echocardiography reports (n = 45,794) were analyzed using GPT-4o-based LLM extraction deployed within a secure health system enclave and were benchmarked against echocardiographic measurements defined in the MIMIC-III dictionary schema. In MIMIC-III, PH was recorded qualitatively (mild/moderate/severe) based on tricuspid regurgitant (TR) jet velocity and then re-coded as present vs. absent. LLM based extraction defined RVD as (1) RV structural abnormality (>= 1 of hypertrophy, dilation, or wall hypo-/akinesis) or (2) RV pressure/volume overload (>= 2 of the following: estimated right atrial pressure > 8 mmHg, TR jet velocity > 2.8 m/s, fractional area change < 35%, tricuspid annular planar systolic excursion < 17 mm, S' < 9.5 cm/s, or E/e' > 14), with PH defined as estimated pulmonary artery systolic pressure > 35 mmHg or qualitative documentation of PH. Results: LLM extraction identified PH in 15,394 (33.6%), RV pressure/volume overload in 14,449 (31.6%), and RV structural abnormalities in 11,955 (26.1%). Co-occurrence was common: overload + structural changes in 9,380 (20.5%), overload + PH in 9,756 (21.3%), structural changes + PH in 6,183 (13.5%), and all three in 5,620 (12.3%). Using the MIMIC-III dictionary schema, PH prevalence was similar (15,371; 33.6%), but RV overload fields were captured less often (pressure overload 1,357 [3.0%], volume overload 1,128 [2.5%], pressure + volume overload 1,093 [2.4%]; any overload field 3,578 [7.8%]), and RV pressure/volume overload with PH was identified in only 731 (1.6%). Conclusions: LLM-based extraction outperforms rules-based schemas for identifying complex disease states not defined by any single variable. By synthesizing multifactorial signals, LLMs can phenotype RVD with higher fidelity and support population-level assessment. Further validation using multimodality imaging, invasive hemodynamics, and clinical outcome data is needed.
ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.
Show abstract
Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.
Gorenshtein, A.; Omar, M.; Jia, E. L.; Adiniaev, Y.; Daniel, O.; Kruskal, J.; Ahmed, M.; Brook, O. R.; Klang, E.; Barash, Y.
Show abstract
Objective: Published P300-speller fusion schemes fix prior trust regardless of trial reliability; we tested whether a reliability estimate improves on it. Methods: We reanalyzed 3,373 archived P300-speller selections from 47 people with ALS (BigP3BCI). A fair, matched-search-space comparison, tuning both a fixed weight and an adaptive policy out-of-fold, was evaluated across 22 evaluable language-model priors up to 46.7B parameters. Two representative priors, GPT-2 and a classical 5-gram, additionally received detailed naive and mechanistic analyses. Results: No prior's 95% CI favored adaptive fusion under the fair comparison, despite unexploited oracle headroom at every scale. Under GPT-2, the naive comparison was significantly worse for adaptive fusion; both anchors converged to a degenerate or near-degenerate fair-comparison solution. For the representative anchors, three further controllers failed to convert that headroom into benefit; the fixed-fused posterior's output probability outperformed the best controller for flagging errors (2.8- to 3.8-fold enrichment). Conclusion: A tuned fixed weight is a difficult-to-beat default across the tested scale range; reliability estimation gave no deployable adaptive advantage. Significance: Adaptive weighting should be validated against a fairly tuned baseline across model families and scales; in this dataset, the fused output's confidence identified high-risk selections better than the tested purpose-built ranker.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Withanage, N. D.; Perera, S.; Athiththan, L.
Show abstract
Background: Lumbar disc herniation, with or without concomitant disc degeneration, is a major cause of lumbar radiculopathy and low back pain, which also a key public musculoskeletal disorder without an exact pathophysiology. Studies have suggested that inflammatory cells and biochemical markers of inflammation also play an important role in lumbar radiculopathy in addition to nerve compression. The aim of the present study was to assess the association of selected circulatory inflammatory markers (CRP, hs-CRP and E-selectin) in patients with lumbar disc herniation without radiological degeneration (LDH) and lumbar disc herniation with radiological degeneration (LDHD). Materials & methods: This case-control study included 208 participants, comprising 104 patients with lumbar disc pathology and 104 controls. Patients were further stratified into LDH (n=67) and LDHD (n=37). Serum CRP, hs-CRP and E-selectin concentrations were measured. Results: Among the patients, 35.6 % presented with LDHD while 64.4 % had only LDH. Significantly increased median hs-CRP (p<0.001) and CRP (p<0.001) were observed in patients groups compared to controls, while CRP showing a consistent independent association across the combined disease (OR=1.68, 95% CI=1.33-2.14, p<0.001), LDHD (OR=1.62, 95% CI=1.16-2.20, p=0.005) and LDH (OR=1.69, 95% CI=1.30-2.20, p<0.001) multivariable models. No significant difference was observed in serum E-selectin between the study groups. Multivariable models incorporating inflammatory and clinical variables demonstrated substantially greater discriminatory performance than individual biomarkers alone. Conclusion: Elevated circulating CRP and hs-CRP concentrations were associated with lumbar disc pathology, with CRP showing a consistent independent association across the combined disease, LDH and LDHD multivariable models, whereas E-selectin showed no significant association. Multivariable models incorporating inflammatory and clinical variables demonstrated greater discriminatory performance than individual biomarkers. These findings support a potential systemic inflammatory component in lumbar disc pathology, although the cross-sectional nature of the measurements does not establish causality or a local inflammatory response within the disc.
Losa, M.; Cotta Ramusino, M.; Gandoglia, I.; Mazzacane, F.; Orso, B.; Lorenzini, L.; Donniaquio, A.; Massa, F.; Sentieri, E.; Gualco, L.; Perini, G.; De Franco, V.; Costa, A.; Bax, F.; Greenberg, S. M.; Kozberg, M. G.; Piazza, F.; Uccelli, A.; Schenone, A.; Del Sette, M.; Farina, L. M.; Roccatagliata, L.; Pardini, M.
Show abstract
Background: The Boston Criteria v2.0 represent the gold standard for diagnosing Cerebral Amyloid Angiopathy (CAA), but their application is currently precluded in mixed small vessel disease (SVD), where deep and lobar hemorrhages coexist. The aims of this study are: (i) to determine which cerebrospinal fluid (CSF) biomarker (A{beta}42, A{beta}40, A{beta}42/40 ratio) is the best candidate to support the CAA diagnosis; (ii) to define a data-driven cut-off, and (iii) to explore if a biomarker-integrated classification significantly improves the phenotypical concordance with the suspected predominant SVD (CAA vs. arteriosclerosis). Methods: We analyzed data from a retrospective multicenter cohort of patients with suspected CAA, defined as probable CAA (Boston criteria v2.0) but allowing deep hemorrhagic lesions, and with available CSF biomarkers. We visually quantified MRI-visible SVD markers (e.g., cerebral microbleeds [CMB], cortical superficial siderosis [cSS], lacunes) and their association with MRI-visible SVD features. We employed a Gaussian Mixture Model (GMM) to identify a data-driven threshold for amyloid positivity (A+). Then, we compared the prevalence of MRI-visible manifestations of SVD between subgroups applying different frameworks, namely the current MRI-based classification (probable CAA vs. mixed SVD) and a CSF biomarker-integrated classification (A+ vs. A-). Results: We enrolled 121 patients (age: 72 [66-77] years; 60% probable CAA, 40% mixed SVD with suspected CAA). The CSF A{beta}42/40 ratio showed a bimodal distribution and consistent associations with all CAA-specific radiological features. The CSF biomarker-integrated reclassification, particularly using the GMM cut-off, significantly improved the distinction between subgroups regarding CAA- and arteriosclerosis-related MRI features (e.g., cSS presence: probable CAA vs. mixed SVD: aOR=2.84 [95%CI 1.27-6.39], p=0.011; A+ vs. A-: aOR=12.68 [95%CI 4.31-37.32], p<0.001; deep lacunes presence: probable CAA vs. mixed SVD: aOR=0.20 [95%CI 0.08-0.50], p<0.001; A+ vs. A-: aOR=0.04 [95%CI 0.01-0.11], p<0.001). Notably, patients classified as A+ never demonstrated more than four deep CMBs. Discussion: A CSF biomarker-integrated classification may improve the classification of CAA compared with the current MRI-based framework. These findings are cohort-specific and would benefit from further validation, especially with a neuropathological reference. Still, these results support a future transition toward an integrated biological-radiological framework, which may refine in vivo CAA diagnosis, particularly in mixed SVD.
LEI, P.; XU, Y.; ZHANG, Y.
Show abstract
Background: The condition of a patient with acute stroke often changes within hours of ICU admission. Prognostic work here targets fixed endpoints predicted from admission data, and trajectory phenotyping assigns one label per patient. We used longitudinal ICU data to identify interpretable dynamic clinical states, characterize transitions between them, and relate the current state to later events. Methods: Retrospective cohort study of 6368 adults with acute stroke in MIMIC IV v3.1. The first 72 h were divided into twelve 6-hour windows, and a hidden Markov model was fitted to 21 neurological, physiological and organ support variables. State number was chosen against criteria fixed before fitting: statistical fit, restart stability, state occupancy and clinical interpretability. Generalized estimating equations related the current state to new mechanical ventilation and vasopressor use within 12 h, and to ICU death within 72 h. Eleven sensitivity analyses assessed the robustness of the state solution. Results: Four states were selected: neurologically preserved-low support, neurological impairment low support, impairment renal dysfunction and impairment-respiratory support (63.3%, 7.8%, 11.8% and 17.1% of windows). Within 72 h, 40.3% of patients changed state at least once, and transitions ran in both directions rather than along a single severity gradient. States were identified without outcome data, yet ICU mortality by last state ranged from 2.9% to 43.9%. Adjusted for age, sex, subtype and Charlson index, the current state remained associated with organ-support escalation and death. State prevalence differed by at most 1.1 percentage points between training and test sets, and 10 of 11 sensitivity analyses gave a stable four-state solution (ARI 0.754 0.955). Conclusions: The early ICU course of acute stroke can be represented as movement among a small number of clinically interpretable states. The representation was reproducible in a held out set and across admission eras, but requires validation in an independent database before any clinical use.
Radoynova, M.; Benouis, M.; schulze, f.; Winter, S.; Bornhauser, M.; Middeke, J. M.; Eckardt, J.-N.
Show abstract
Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignettes. Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%) and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs. Between model generations, the largest improvements in accuracy were seen for open-weight models whereas proprietary models showed only marginal gains. In error analysis, top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases. Frontier LLMs exhibit substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets. Yet, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is paramount.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Buzzanca, G.; Pala, C.; He, J.; Hofstraat-Boersma, R.; Tammaro, A.; van Midden, D.; Buelow, R.; Hoelscher, D. L.; Muehlfeld, A. S.; Koeller, m.; Kozakowski, N.; Boehmig, G.; Halloran, P. F.; van der Helm, D.; Meziyerh, S.; Venhuizen, J.-H.; Haitjema, S.; Dijkstra, J.; Hilbrands, L. B.; Steenbergen, E. J.; van Zuilen, A. D.; Nurmohamed, A. S.; Bemelman, F. J.; Bruns, I. B.; Callegaro, G.; van de Water, B.; Pieters, T. T.; Breimer, G. E.; Rossi, G. M.; Fiaccadori, E.; Maggiore, U.; Roelofs, J. J. T. H.; Testa, F.; Fontana, F.; Abiola, A. A.; Delsante, M.; Corthals, G. L.; Peters-Sengers, H.; Ngu
Show abstract
Accurate, reproducible interpretation of kidney allograft biopsies is critical for diagnosis of graft injury to guide prognosis and management. The international Banff classification is a consensus diagnostic system based on semiquantitative histological lesion scoring on either extent or severity of kidney transplant biopsies. However, pathologist scoring is limited by substantial interobserver variability, constrained scalability, and the inherent nature of the scoring system itself. Here we present BanffNET, a weakly supervised, probabilistic deep learning framework that combines self-supervised feature extraction with a novel Bayesian multiple-instance learning framework to predict (continuously) the full spectrum of Banff lesion scores directly from whole-slide images (WSIs). Using lesion-specific aggregation functions tailored to localized (modeling lesion severity) and diffuse pathologies (modeling lesion extent), BanffNET generates interpretable, patch-level probability maps and calibrated slide-level scores. BanffNET's performance was assessed relative to consensus, biological correlates of rejection and clinical outcome, demonstrating superior consistency, transportability and generalization. Trained on 7,249 WSIs from three cohorts, BanffNET demonstrates consistent performance on 11,028 WSIs across five external test sets, performing on par or exceeding expert consensus across lesions. BanffNET scores align more closely than pathologist Banff scores with molecular profiles of rejection, offering a transparent, biologically grounded framework for computational pathology with relevance beyond transplantation.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.
Yano, Y.; Shintani, E.; Arita, S.; Ashine, R.; Iinuma, N.; Mori, H.; Fujibayashi, K.; Yamada, Y.; Saita, M.; Nakashima, N.; Itoh, H.; Nangaku, M.; Ohashi, M.; Daida, H.; Arai, H.; Naito, T.
Show abstract
The widespread adoption of clinical large language models (LLMs) introduces significant risks of automation bias, premature closure, and clinician deskilling. Current interpretability paradigms, including latent space trajectories, Concept Activation Vectors, and Concept Bottleneck Models, suffer from topological stagnation, metric distortion, and epistemic occlusion, frequently masking intermediate diagnostic uncertainty behind falsely confident outputs. To address these structural vulnerabilities, this paper introduces a novel closed-loop, multi-agent framework designed to quantify and visualize dynamic epistemic uncertainty in clinical LLM reasoning. By coupling predictive Shannon entropy with non-linear Isometric Feature Mapping (ISOMAP), the architecture projects high-dimensional inference state vectors onto a calibrated two-dimensional latent space, thereby assigning a quantifiable thermodynamic energy state to the reasoning path to track diagnostic velocity, cognitive momentum, and trajectory efficiency across sequential diagnostic rounds. Pilot validation across representative emergency medicine scenarios demonstrated distinct topological and information-theoretic behaviors: unconfounded cases (cerebellar infarction) exhibited smooth geodesic progression toward the ground truth alongside monotonic Shannon entropy decay from 2.15 to 1.74; noisy environments with ambiguous findings (spontaneous pneumothorax) suffered from trajectory wandering, local minimum traps, and high sustained entropy (~2.41) due to insufficient repulsive weighting for negative evidence; and triage-conflicted cases (acute cholangitis) achieved precise geometric proximity to the true node but experienced top-1 rank stagnation because the model conflated acute severity triage (sepsis) with anatomical etiology. By rendering machine hesitation and cognitive divergence visually auditable before final diagnostic crystallization, this geometric-information framework enables dynamic trust calibration and human-AI co-regulation at the point of care while establishing a clear mathematical foundation for future architectural interventions, such as dual-channel safety decoupling and non-linear repulsive weighting. Moving forward, validating these architectural enhancements across large-scale electronic health record databases and prospective clinical trials will be essential to realize its full clinical utility, establishing a foundational blueprint for safe, transparent, and cognitively synergistic AI integration in future medical practice. By rendering the LLM's reasoning process visually auditable, this framework lays the groundwork for capturing and externalizing the clinician's own cognitive patterns within the AI, forming a coupled system. This enables the explicit visualization of cognitive gaps between physician hypotheses and AI inferences, transforming the interaction from simple answer-checking into a dynamic learning process for both human and machine that prevents diagnostic oversight. Ultimately, because the responsibility for final clinical decision-making remains with the human practitioner, this framework serves as a vital decision-support mechanism. Moving forward, validating these architectural enhancements across large-scale electronic health record databases and prospective clinical trials will be essential to realize its full clinical utility, establishing a foundational blueprint for safe, transparent, and cognitively synergistic AI integration in future medical practice.
Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.
Show abstract
Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.
Iliadis, I.; Heitland, I.; Hoeper, K.; Witte, T.; Kahl, K. G.; Stapel, B.; Meyer-Olson, D.
Show abstract
Objective: The Brief-cope questionnaire explore coping behavior. However, the underlying factor structure remains a subject of ongoing debate. Exploratory factor analyses (EFA) conducted across different populations have identified factor solutions ranging from two to fourteen factors. As of yet, the underlying factor structure of the Brief-cope has not been investigated in patients with seropositive rheumatoid arthritis (RA). Therefore, the aim of this study was to explore the underlying factor structure of the Brief-cope in a German population of seropositive RA. Methods: 216 outpatients with seropositive RA completed the Brief-cope. An EFA with principal axis factoring and Promax rotation was conducted. Results: EFA indicated a five-factor solution. The five-factor solution explained 51.95% of variance. The identified factors were: (1) problem-focused coping (Cronbach's = .851), (2) emotion-focused coping ( = .754), (3) maladaptive coping ( = .747), (4) religious coping ( = .851), and (5) substance-use coping ( = .869). Conclusion: A five-factor solution provided the most appropriate representation of the underlying factor structure of the Brief-cope in patients with seropositive RA. This factor structure may serve as a suitable basis for future analyses of Brief-cope data in comparable RA populations.
Bou Dagher, L.; Han, Z.; Zhou, S.; Fülöp, T.; Desroches, M.; Rodrigues, S.
Show abstract
Alzheimer's disease is characterized by the accumulation and aggregation of amyloid-{beta}(A{beta}), but the molecular mechanisms linking environmental and infectious factors to A$\beta$ conformational changes remain incompletely understood. Herpes simplex virus type 1 (HSV-1) has been proposed as a potential contributor to AD pathology, and interactions between the viral glycoprotein B (gB) and A$\beta$ may influence the conformational behaviour of the peptide. Molecular dynamics (MD) simulations provide atomic-scale information on such interactions, but conventional structural descriptors may not fully capture changes in the organization of residue interaction networks. Here, we introduce a graph-geometric framework based on Forman-Ricci curvature to characterize the evolution of residue interaction networks during MD simulations. Each simulation frame is represented as a residue interaction graph based on C--C contacts, and residue-wise curvature profiles are analysed across time. We apply the framework to A{beta}1-42 in isolation and in complex with HSV-1 gB. Conventional MD analyses indicate stable association of the simulated complex, favourable interaction energetics, and conformational changes in A{beta}, including a transition from -helical structure toward {beta}-turn-rich conformations over the simulated timescale. Forman-Ricci curvature reveals pronounced and spatially localized remodelling of the A{beta} residue interaction network in the complex, with the strongest changes concentrated in the C-terminal region. These regions also exhibit reduced temporal curvature fluctuations and progressively distinct geometric behaviour throughout the simulation. Hierarchical clustering further identifies cooperative groups of residues with coordinated curvature dynamics, including a prominent C-terminal domain. Together, these results demonstrate that Forman-Ricci curvature provides a complementary description of biomolecular dynamics by capturing changes in the geometric organization of residue interaction networks that are not directly represented by conventional structural descriptors. The framework provides a general computational approach for studying network-level structural remodelling in protein molecular dynamics and offers a quantitative perspective on the conformational consequences of HSV-1 gB--A{beta} association.
Sadia, H.; Doyon, N.; Duchesne, S.
Show abstract
Background Understanding the mechanisms underlying brain aging and age-related pathological changes is essential for advancing brain health research. Our group previously developed a mechanistic mathematical model of healthy brain, Chamberland et al. (2024) that integrates key biological processes involved in normal aging, from which Alzheimer's disease (AD) related changes may emerge naturally. Objectives To characterize and validate this brain model by evaluating its sensitivity, calibrating its parameters, and assessing generalizability in independent populations. Methods The model represents the evolution of key biological processes associated with brain aging, including amyloid beta (A{beta}), tau pathologies, neuroinflammation, and neuronal death. After identifying the 30 most influential parameters, we calibrated the model using cognitively normal (CN) participants from the AD Neuroimaging Initiative (ADNI) database (n = 211) by minimizing a loss function composed of three outcomes (AB) plaques, tau tangles, and neuronal density). The calibrated model was then applied to the UK Biobank cohort (n = 35,899) of normal controls (aged 44-82 years). The effects of sex and APOE were evaluated using stratified simulations. Results Parameter calibration significantly reduced the prediction errors for A{beta} and tau. Neuronal density predictions showed strong agreement in the UK Biobank cohort. The variance decomposition identified APOE status as a major contributor to variability in A{beta}. Conclusion Our validated brain health model links mechanistic pathways with population data and reproduces neuronal density patterns in an independent cohort. These findings support its use as a framework for studying brain aging and investigating how Alzheimer's disease related pathological changes may emerge with aging.